Papers with automated evaluation metrics

16 papers
Generating Text through Adversarial Training Using Skip-Thought Vectors (N19-3)

Copied to clipboard

Challenge: Existing approaches to use word embeddings for text generation have been limited.
Approach: They propose to use GANs with word embeddings to reproduce writing style in text . they use a sentence embeddable vector to model people's way of expression .
Outcome: The proposed model outperforms baseline text generation networks across several metrics including BLEU-n, METEOR and ROUGE.
Vis-Eval Metric Viewer: A Visualisation Tool for Inspecting and Evaluating Metric Scores of Machine Translation Output (N18-5)

Copied to clipboard

Challenge: Many metrics have been proposed for Machine Translation (MT) that compare system translations against human references.
Approach: They propose to use BLEU and METEOR to evaluate machine translations against human translations.
Outcome: VisEval Metric Viewer (VEMV) provides visualisation of multiple evaluation scores so they can be easily interpreted by a user.
Faithful Low-Resource Data-to-Text Generation through Cycle Training (2023.acl-long)

Copied to clipboard

Challenge: Methods to generate text from structured data have advanced significantly in recent years, but can fail to produce output faithful to the input data, especially on out-of-domain data.
Approach: They evaluate the effectiveness of cycle training by using two models which are inverses of each other to generate text from structured data and one which generates the structured data from natural language text.
Outcome: The proposed approach achieves nearly the same performance as fully supervised approaches on the WebNLG, E2E, WTQ, and WSQL datasets.
Towards Better Evaluation for Generated Patent Claims (2025.acl-long)

Copied to clipboard

Challenge: Existing studies highlight inconsistencies between automated evaluation metrics and human expert assessments for patent claims.
Approach: They propose a multi-dimensional evaluation method specifically designed for patent claims that incorporates features annotated by patent experts.
Outcome: The proposed method achieves highest correlation with human expert evaluations across all assessment criteria across all tested metrics.
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Current evaluation methods do not capture complex attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence.
Approach: They propose to use a dataset and framework to evaluate counterspeech quality across four dimensions: contextual relevance, aggressiveness, argument-coherence, and suitability.
Outcome: The proposed method outperforms ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation.
COSMic: A Coherence-Aware Generation Metric for Image Descriptions (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to evaluate captions have limited learning of their output . previous methods focused on n-gram measures of similarity to reference output based on a ngram of similarities to the output metric.
Approach: They propose a first discourse-aware learned generation metric for evaluating image descriptions.
Outcome: The proposed metric predicts human ratings of captions on out-of-domain images.
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics cannot fairly evaluate the outputs of RAG models during training and evaluation.
Approach: They propose a method which prompts LLMs to generate different judgments based on various combinations of judgment dimensions and utilizes the judge-consistency to evaluate these judgments.
Outcome: The proposed method generates more accurate evaluations for RAG models across different RAG model and datasets.
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding (2025.acl-long)

Copied to clipboard

Challenge: Understanding domain-specific theorems requires more than text-based reasoning . current evaluations of theoretical models are based on textual cues .
Approach: They propose an agentic approach for generating long-form theorem explanation videos using Manim animations.
Outcome: The proposed agent generates long-form theorem explanation videos using Manim animations . the agent achieves a success rate of 93.8% and an overall score of 0.77 .
Evaluation of Thematic Coherence in Microblogs (2021.acl-long)

Copied to clipboard

Challenge: Recent work on grouping together views about tweets expressing opinions about the same entities has been criticized for their lack of thematic coherence.
Approach: They propose to use a corpus of microblogs representing opinions about the same topics within the same time window to evaluate thematic coherence.
Outcome: The proposed method outperforms surface level metrics, topic model coherence and text generation metrics (TGMs) but is not as reliable as TGMs due to being less sensitive to time windows.
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations (2023.acl-long)

Copied to clipboard

Challenge: Prior work has shown that models may exploit shortcuts that are difficult to detect using standard n-gram similarity metrics such as ROUGE.
Approach: They propose to use human-assessed summary quality facets and pairwise preferences to improve MDS evaluation methods.
Outcome: The proposed methods improve the quality of literature review summarization models . they use human-assessed summary quality facets and pairwise preferences .
“A good pun is its own reword”: Can Large Language Models Understand Puns? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies on the understanding of puns in large language models (LLMs) have not explored the use of pun in creative writing and humor creation.
Approach: They propose to use pun recognition, explanation and generation tasks to evaluate the capabilities of large language models (LLMs) they adopt automated evaluation metrics from prior research and introduce new evaluation methods and metrics that align more closely with human cognition.
Outcome: The proposed methods align more closely with human cognition than previous evaluation metrics.
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing "LLM-as-a-judge" evaluation frameworks are limited by persona descriptions and are not generalizable to other tasks.
Approach: They propose a framework that can automatically construct multiple evaluator personas with distinct dimensions from relevant text documents and instantiate LLM agents with the persona.
Outcome: The proposed framework can believably simulate human evaluators . it extracts stakeholders' diverse perspectives from the provided research papers and constructs personas for the agents .
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations (2025.findings-acl)

Copied to clipboard

Challenge: Language models are widely used in education, yet their ability to tailor responses to learners with varied informational needs and knowledge backgrounds remains under-explored.
Approach: They conduct two extensive human studies to assess the utility of language model-generated explanatory answers (explanations) on a benchmark of 13.4K "Why" questions.
Outcome: The proposed model explanations match learners' educational backgrounds only 50% of the time, compared to 79% for lay explanations.
TaxoAlign: Scholarly Taxonomy Generation Using Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for taxonomy generation do not compare structure of generated surveys with those written by human experts.
Approach: They propose a method that bridges the gap between human-generated and automatically-created taxonomies.
Outcome: The proposed method surpasses baselines on CS-TaxoBench on nearly all metrics.
Would You Like to Make a Donation? A Dialogue System to Persuade You to Donate (2024.lrec-main)

Copied to clipboard

Challenge: Persuasive automated dialogue systems are a popular way to influence people's behavior and decision making.
Approach: They propose to use a context-aware persuasion strategy selection module to persult users . they also propose a persuasiveness prediction model to automatically evaluate the persuasiveness of generated text.
Outcome: The proposed system can achieve better performance on several automated evaluation metrics than baseline models.
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) produce incomplete or selectively omit key information . omissions of key information or misrepresentation of conflicting evidence can cause harm .
Approach: They propose a method that decomposes texts into atomic statements and uses natural language inference to identify missing facts and a Q A-based metric that extracts question-answer pairs and compares responses across sources.
Outcome: The proposed evaluation metrics show they perform better than more complex metrics, but at a cost.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations